Papers with multimodal GUI agent framework
INREACT: An Inspire-Then-Reinforce Training Framework For Multimodal GUI Agent (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing multimodal large language models struggle with precise localization of small elements. |
| Approach: | They propose a multimodal GUI agent framework that unifies observing, thinking, and acting for precise and interpretable decision-making. |
| Outcome: | The proposed framework unifies observing, thinking, and acting for precise and interpretable decision-making. |